Papers with open-source models
Copied to clipboard
| Challenge: | MERaLiON-AudioLLM is the first general-purpose audio-based large language model for multitask learning. |
| Approach: | They introduce MERaLiON-AudioLLM, a general-purpose audio-based large language model for multitask learning with a focus on Singlish understanding. |
| Outcome: | The proposed model exhibits strong generalization across a diverse set of tasks . it is a leading solution for region-specific AI applications. |
Copied to clipboard
| Challenge: | Full-duplex speech agents are often half-duplice, alternating turns between user and system. |
| Approach: | They propose a streaming framework that integrates with an examiner that enforces staged goals under two pacing setups. |
| Outcome: | The framework reports fluency, multi-turn instruction following, and task-specific competence. |
Copied to clipboard
| Challenge: | Diagram-grounded geometry problem solving is critical for multimodal large language models, but the benefits of multi-agent design over single-aggent remain unclear. |
| Approach: | They compare diagram-grounded geometry problem solving to four visual math benchmarks . they found that multi-agent pipelines provide clear benefits for open-source models . |
| Outcome: | Theorem-based solvers and architectural refinements improve performance on four visual math benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have sparked interest in creating autonomous agents. |
| Approach: | They propose a framework that jointly optimizes both task-planning and self-reflective evolution capabilities in language agents. |
| Outcome: | The proposed framework improves task planning and self-reflective evolution capabilities in language agents. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been shown to be effective in complex language tasks, but their potential to perpetuate biases poses significant concerns. |
| Approach: | They propose a new framework employing Direct Preference Optimization to mitigate biases in LLMs. |
| Outcome: | The proposed model outperforms the baseline model on almost all bias benchmarks and achieves better performance than open-source models. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have enabled the development of advanced agentic systems that can integrate various tools and APIs to fulfill user queries. |
| Approach: | They propose an end-to-end framework for training and deploying task-specific small language model agents capable of function calling for driving agentic systems at the edge. |
| Outcome: | The proposed model outperforms existing models by reducing the input prompt length and quantizing the inference speed. |
Copied to clipboard
| Challenge: | Adapting general multimodal large language models to specific domains is important for practical applications. |
| Approach: | They investigate domain adaptation of multimodal large language models via post-training . they develop a generate-then-filter pipeline that curates diverse visual instruction tasks . |
| Outcome: | The proposed model outperforms existing models in domain adaptation by combining data from open-source models with training pipelines. |
Copied to clipboard
| Challenge: | open-source application for real-time multilingual bi-directional translation between spoken and signed languages. |
| Approach: | They present an open-source application for real-time multilingual bi-directional translation between spoken and signed languages. |
| Outcome: | The open-source sign.mt application aims to address the communication divide between the hearing and the deaf. |
Copied to clipboard
| Challenge: | Positional biases in large language models hinder their ability to process long inputs. |
| Approach: | They propose a benchmark to assess positional bias in large language models involving multiple pieces of relevant information. |
| Outcome: | The proposed benchmark assesses the performance of long-context language models by examining their models with different input lengths and tasks. |
Copied to clipboard
| Challenge: | Existing large language models favor high-resource languages, such as English, at the expense of low-resourced and regional languages. |
| Approach: | They propose a series of language models that specifically focuses on Southeast Asian languages. |
| Outcome: | SeaLLM models outperform ChatGPT-3.5 in non-Latin languages by large margins . linguistic disparity impedes access to state-of-the-art AI technologies for non-English-speaking populations . |
Copied to clipboard
| Challenge: | Membership inference attacks are a canonical way to assess a machine learning model’s privacy properties. |
| Approach: | They propose a framework for principled evaluation of membership inference attacks against large language models by leveraging the insight that training data before and after a fixed point during training are drawn from the same distribution. |
| Outcome: | The proposed framework can be used to evaluate membership inference attacks against large language models. |
Copied to clipboard
| Challenge: | a study examines how to build meeting summarization systems using large language models . closed-source models are generally better in terms of performance, but open-source ones are more advantageous for industrial use . |
| Approach: | They compare closed-source and open-source meeting summarization models for real-world use . they find that closed-sourced models are generally better in terms of performance . however, smaller open-sourced LLMs could still achieve comparable performance if they are open . |
| Outcome: | The proposed model is more efficient for industrial use than closed-source models due to privacy concerns and high cost. |
Copied to clipboard
| Challenge: | Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). |
| Approach: | They propose a framework for detecting pro-bigorexia content on TikTok using an expert-annotated multimodal benchmark dataset of over 2,200 Tiktok videos labeled by clinical psychiatrists. |
| Outcome: | The proposed framework improves on fine-grained subcategories while commercial models achieve the highest accuracy on primary categories. |
Copied to clipboard
| Challenge: | *Dialz* is a Python library for advancing research on steering vectors for open-source LMs. |
| Approach: | They propose a Python library for advancing research on steering vectors for open-source LMs. |
| Outcome: | The proposed method reduces harmful outputs and provides insights into model behaviour across different layers. |
Copied to clipboard
| Challenge: | Existing models for learning large language models are expensive and difficult to build and fine-tune. |
| Approach: | They propose a family of data augmentation models to improve model fine-tuning efficiency . they leverage powerful LLMs to expand, refine and re-write instructions and responses . |
| Outcome: | The proposed models improve the efficiency of model fine-tuning by leveraging small datasets and quality assessment techniques. |
Copied to clipboard
| Challenge: | Existing evaluations of multimodal abductive reasoning are limited to static, single-agent tasks. |
| Approach: | They propose a multiagent evaluation suite that deconstructs the current evaluations of multimodal abductive reasoning in vision–language models. |
| Outcome: | The evaluation suite is based on two core components: DixitArena and DixitsBench. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to assess large language models’ format-following proficiency adequately. |
| Approach: | They propose a benchmark to evaluate large language models' ability to follow complex, domain-specific formats. |
| Outcome: | The proposed framework evaluates large language models' ability to follow complex, domain-specific formats across open-source and closed-source models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have evolved into interactive agents capable of planning, tool use, and task execution across various tasks. |
| Approach: | They propose a platform that leverages large language models to generate agent-tuning data for fine-tuneing smaller, specialized models. |
| Outcome: | MIMIR enables large models to simulate various roles and create interaction data, which can then be used to fine-tune open-source models like LLaMA2. |
Copied to clipboard
| Challenge: | Unlike existing tools, our system addresses the ambiguity of vague, multi-line queries, setting a new benchmark in data storytelling by tackling complexities no existing system comprehensively handles. |
| Approach: | They propose a system that processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories. |
| Outcome: | The proposed system processes and interprets vague, open-ended, and multi-line complex queries, transforming them into coherent, actionable data stories. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are a promising way to bridge the gap between patient health literacy and access to care. |
| Approach: | They evaluate a range of open- and closed-source LLMs on a MeDiSumQA dataset . they propose a lightweight multi-agent framework for patient-oriented medical question answering . |
| Outcome: | The proposed model achieves lower FKGL than zero-shot GPT-5 and highest simplification quality among all models. |
Copied to clipboard
| Challenge: | Large Language Models excel in code generation benchmarks, but these benchmarks focus on single-file scenarios with constrained context scope. |
| Approach: | They propose an open-source framework to effectively resolve GitHub issues using a code file retrieval module and a model-based code editing module. |
| Outcome: | The proposed approach achieves state-of-the-art performance on two GitHub benchmarks. |
Copied to clipboard
| Challenge: | a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models. |
| Approach: | They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models . |
| Outcome: | The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions . |
Copied to clipboard
| Challenge: | proprietary large language models (LLMs) have demonstrated impressive code generation performance. |
| Approach: | They propose an adaptive module-based model that refines the direct response distillation process by modular decomposition and adaptive response evolution. |
| Outcome: | The proposed framework outperforms baseline model and code generation methods on three popular benchmarks. |
Copied to clipboard
| Challenge: | Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. |
| Approach: | They propose a Monte Carlo Tree Search (MCTS) based Prompt Auto-generation (MPA) method to generate adversarial suffixes for valid jailbreak attacks. |
| Outcome: | The proposed method outperforms existing methods on open-source and closed-source models and shows that it can generate harmful responses. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have significant impact on various industries and societal functions due to advanced instruction-following capabilities. |
| Approach: | They developed and tested novel attack strategies on popular LLMs to expose their vulnerabilities in generating harmful content. |
| Outcome: | The proposed attacks achieved an ASR of 100% on open-source models, including Meta’s Llama-3.2, Google’s Gemma-2, Mistral’s Mistral-NeMo, Falcon’s Falcon-mamba, Apple’s DCLM, Microsoft’s Phi3, and Qwen’s Qwend2.5, among others. |
Copied to clipboard
| Challenge: | Chartered Financial Analyst (CFA) program is widely recognized globally . study compares state-of-the-art large language models with open-source models . proprietary models pass levels I and II, but fail at level III due to essay questions . |
| Approach: | They benchmark five leading proprietary models and eight open-source models on mock CFA exams to provide an overview of their financial analysis capabilities. |
| Outcome: | The models on the mock CFA exams pass the highest scores, but fail at the lowest levels due to essay questions. |
Copied to clipboard
| Challenge: | Recent Long-Context Language Models (LCLMs) do not capture how evidence should be connected . a new framework that integrates thought templates into LCLM frameworks is proving useful . |
| Approach: | They propose a framework that iteratively refines reusable reasoning patterns derived from prior problem solving to improve their templates. |
| Outcome: | The proposed framework outperforms baselines on knowledge-intensive multi-hop reasoning benchmarks and practical scenarios without retrieval. |
Copied to clipboard
| Challenge: | Existing benchmarks for Table Information Seeking (TabIS) are lacking in reliable evaluation. |
| Approach: | They propose a benchmark to evaluate the table information seeking abilities of large language models . they use a single-choice question format instead of a text-based evaluation . |
| Outcome: | The proposed benchmark is more reliable than existing models and is available online. |
Copied to clipboard
| Challenge: | evaluating the generalisability of Transformers to out-of-distribution mathematical reasoning problems is a challenge for many open-source models. |
| Approach: | They propose a method for generating and perturbing detailed derivations of equations at scale, aided by a symbolic engine, and compare their results to sequence classification tasks. |
| Outcome: | The proposed framework outperforms GPT-4, GPT-3.5 and a canon of fine-tuned BERT models in classification tasks . perturbations to input reasoning can reduce their performance by up to 80 F1 points . |
Copied to clipboard
| Challenge: | Recent advances in LLMs have significantly improved mathematical problem-solving, with models like GPT-4 achieving human-level performance. |
| Approach: | They propose a bilingual English-Korean dataset enriched with teacher solutions, student solutions, and annotations marking students’ initial errors. |
| Outcome: | The proposed model achieves high agreement with human judgments and lower latency and resource usage than commercial APIs, demonstrating strong computational efficiency. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate hallucinations when handling unfamiliar information. |
| Approach: | They propose a Chinese benchmark to evaluate large language models' knowledge and reasoning capabilities in medication tasks. |
| Outcome: | The proposed benchmark evaluates models in indication, dosage and administration, contraindicated population, mechanisms of action, drug recommendation, and drug interaction across six datasets. |
Copied to clipboard
| Challenge: | Large language models are trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. |
| Approach: | They propose a set of requirements that practical contamination detection methods should follow to effectively detect benchmark contamination in large language models. |
| Outcome: | The proposed method detects whether the model is significantly more confident under the original benchmark. |
Copied to clipboard
| Challenge: | Existing benchmarks for mathematical reasoning are becoming less effective due to performance saturation. |
| Approach: | They propose to use a mathematical reasoning benchmark with Olympiad difficulty to evaluate top-tier LLMs. |
| Outcome: | The proposed benchmarks are cross-validated by experts to meet IMO difficulty standards and entirely original problems to prevent performance leakages from data memorization. |
Copied to clipboard
| Challenge: | Recent advances in text-to-SQL generation rely on large closed-source models that present challenges in accessibility, privacy, and latency. |
| Approach: | They propose to use open-source text-to-SQL models to critique SQL queries . their method evaluates multiple outputs simultaneously and is competitive with larger models . |
| Outcome: | The proposed method achieves state-of-the-art performance compared to open-source models while remaining competitive with larger models at a much lower cost. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) have unlocked many complex use cases that require Multi-Modal (MM) understanding and MM generation. |
| Approach: | They propose a plug-and-play technique that adds relevant retrieved information to prompts as few-shot examples during inference. |
| Outcome: | The proposed method significantly improves the output quality of large vision language models when input prompts are augmented with relevant information retrieved by Vision-Language retrievers like UniRAG. |
Copied to clipboard
| Challenge: | Open-source LLMs often depend on large proprietary models, which introduce serious privacy concerns. |
| Approach: | They propose a plug-and-play framework that improves SQL generation for smaller LLMs . they propose to apply question decomposition at the schema linking stage rather than during SQL generation . |
| Outcome: | The proposed framework improves schema linking recall by 25.1% and execution accuracy by 8.2% on the BIRD benchmark. |
Copied to clipboard
| Challenge: | Existing research focuses on enhancing large language models through scaling laws or fine-tuning strategies, but ignores the potential of using agent paradigms to compensate for the inherent weaknesses of small models. |
| Approach: | They propose to use structured agent frameworks to improve effectiveness over direct prompting . they also propose to employ routing-based multi-agent systems with collaborative capabilities . |
| Outcome: | The proposed model significantly outperforms direct prompting with single-agent systems . the proposed model is more reliable and cost-effective than other models . |
Copied to clipboard
| Challenge: | Recent advances in natural language processing and computer vision have made it possible to translate images with text in one language into equivalent images displaying that text translated into another language. |
| Approach: | They propose an all-encompassing framework for the task–In-Image Machine Translation (IIMT) that incorporates contextual cues from both textual and visual elements during translation. |
| Outcome: | The proposed framework can be constructed using open-source models and requires no training, making it highly accessible and expandable. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, but their proficiency in mathematical reasoning remains a key challenge. |
| Approach: | They propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |
| Outcome: | The proposed framework evaluates LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |
Copied to clipboard
| Challenge: | Unsupervised multitask pre-training has been the key to the success of language models (LMs) however, scaling it in the post-training stage trends towards better generalization. |
| Approach: | They propose a framework that augments massive raw corpora with instruction-response pairs to pre-train LMs. |
| Outcome: | The proposed framework augments massive raw corpora with instruction-response pairs to pre-train LMs. |
Copied to clipboard
| Challenge: | Existing models have demonstrated outstanding capabilities in mathematical reasoning, but there is a performance gap between open-source models and closed-source ones. |
| Approach: | They propose a method for generating diverse and reliable math problems by leveraging the ground-truth solutions of the seed data. |
| Outcome: | The proposed model outperforms open-source models across five representative mathematical reasoning datasets. |
Copied to clipboard
| Challenge: | Existing factuality metrics cannot evaluate paragraphs with ambiguous entities, authors show . |
| Approach: | They propose a new metric to evaluate the factuality of long-form generations from large language models. |
| Outcome: | The proposed metric can assess the factuality of people biographies with entity ambiguity better than FActScore. |
Copied to clipboard
| Challenge: | Existing research has studied privacy in LLM training data memorization, but it does not prevent users from disclosing PII at inference time. |
| Approach: | They propose a task for chaining API-based and local LLMs that uses public data to construct a benchmark that contains personally identifiable information (PII) |
| Outcome: | The proposed model maintains high response quality for 85.5% of user queries while restricting privacy leakage to only 7.5%. |
Copied to clipboard
| Challenge: | Evaluating the performance of LLMs in multi-turn interactions presents significant challenges due to the complexity and variability of user behavior. |
| Approach: | They propose a benchmark framework for assessing LLMs’ function-calling capabilities in multi-turn dialogues. |
| Outcome: | The proposed framework is based on a dataset derived from popular mobile apps and anonymized user logs. |
Copied to clipboard
| Challenge: | a recent study validates the effectiveness of chat language models by fine-tuning instruction data. |
| Approach: | They propose to use a large-scale dataset of instructional conversations to fine-tune a conversational model on instruction data. |
| Outcome: | The proposed model outperforms open-source models in key metrics including scale, average length, diversity, coherence, etc. |
Copied to clipboard
| Challenge: | TurkingBench is a benchmark consisting of tasks presented as web pages with textual instructions and multi-modal contexts. |
| Approach: | They propose to use HTML pages to perform various annotation tasks on crowdsourcing platforms. |
| Outcome: | The proposed model outperforms other models on the TurkingBench benchmark. |
Copied to clipboard
| Challenge: | Existing models struggle with complex queries, especially multi-table joins and reasoning. |
| Approach: | They propose to build a model with synthetic training samples and a structure-aware curriculum learning framework for enhancing SQL generation. |
| Outcome: | The proposed model improves on the existing model on the Spider and Bird benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced rapidly from conversational problem solving to addressing real-world tasks involving tool use, such as software engineering (SWE). |
| Approach: | They propose to build an LLM-based software engineering agent that synthesizes test cases and scales up agent trajectories to build training data. |
| Outcome: | The proposed model outperforms state-of-the-art models on the SWE-bench-Verified benchmark. |
Copied to clipboard
| Challenge: | Existing long-text evaluation benchmarks, such as L-Eval and LongBench, focus on QA and summarization tasks. |
| Approach: | They propose a length-adaptable benchmark for evaluating the long-context understanding of large language models. |
| Outcome: | The proposed benchmarks do not cover ultralong settings (100k+ tokens) and are difficult to evaluate across different length ranges. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have proven to be highly effective in addressing a wide range of complex tasks. |
| Approach: | They propose a method that asks teachers to identify and explain student’s mistakes and then asks them to provide customized instruction learning data. |
| Outcome: | The proposed method reduces the chance of teachers guessing incorrectly with flawed rationales, improving instructional data quality. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on static single-step calculations with explicit instructions. |
| Approach: | They propose a benchmark for evaluating medical calculators in realistic scenarios . they use 118 scenario tasks across 4 clinical domains to evaluate medical calculator performance . |
| Outcome: | The first benchmark for evaluating medical calculators in realistic scenarios is released . it features 118 scenario tasks across 4 clinical domains and is based on a model context protocol integration. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated exceptional coding capabilities, but their debugging capabilities remain relatively unexplored. |
| Approach: | They propose a debugging benchmark consisting of 4,253 LLMs with four major bug categories and 18 minor types in C++, Java, and Python. |
| Outcome: | The proposed benchmark covers four major bug categories and 18 minor types in C++, Java, and Python. |
Copied to clipboard
| Challenge: | Existing work on extrapolating positional embedding (RoPE) has limited results in the application of long context language models. |
| Approach: | They propose a set of parameterized extrapolation functions applied to each layer and attention head to adaptively adjust its extrapolations scales. |
| Outcome: | The proposed model achieves stable extrapolation on 64k contexts by training on 16k length text. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models and their multimodal counterparts have shown significant performance disparities across different languages and cultural contexts. |
| Approach: | They propose to evaluate LLMs on diverse vision-language tasks within a multilingual and multicultural context using M5 benchmark. |
| Outcome: | The proposed benchmarks highlight task-agnostic performance disparities between languages and cultural contexts. |
Copied to clipboard
| Challenge: | Large language model context lengths have increased by at least 1000 in the past seven years . however, longer contexts pose challenges to system instruction following . |
| Approach: | They propose to formalize verifiable instructions to evaluate model compliance . they implement and evaluate six mitigation strategies to enhance instruction compliance in extended contexts. |
| Outcome: | The proposed model performs better in long contexts than in natural language models. |
Copied to clipboard
| Challenge: | Current research on long-form context in Large Language Models (LLMs) focuses on understanding of long-contexts, but the open-ended Long Text Generation (Open-LTG) remains underexplored. |
| Approach: | They propose a method that uses data synthesis and a reward signal to enhance model performance. |
| Outcome: | The proposed method outperforms GPT-4-Turbo and improves performance by 20% on the Open-LTG task. |
Copied to clipboard
| Challenge: | Current medical benchmarks have limitations in question design, data sources and evaluation methods. |
| Approach: | They propose a new benchmark covering five core medical areas . it includes 2,996 questions created from real-world electronic health records . |
| Outcome: | The proposed model covers five core medical areas and includes 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios. |
Copied to clipboard
| Challenge: | Long-context models (LCMs) have seen remarkable advancements in recent years, facilitating tasks like long-document QA. |
| Approach: | They propose an out-of-the-box suite that can assess both generation quality and fidelity in long-context understanding tasks. |
| Outcome: | The proposed suite can assess both generation quality and fidelity in long-context understanding tasks. |
Copied to clipboard
| Challenge: | Watermarking is a key technique for detecting AI-generated text. |
| Approach: | They propose a method to selectively smooth watermarks by leveraging the relationship between the model’s confidence and detectability. |
| Outcome: | The proposed method selectively smoothes watermark traces while preserving text quality. |
Copied to clipboard
| Challenge: | Existing fingerprinting methods for large vision-language models rely on backdoors to elicit abnormal outputs, but direct distortion of the model’s original outputs compromises modality alignment and degrades multimodal capabilities. |
| Approach: | They propose to embed a robust fingerprint while preserving the original normal outputs of the model. |
| Outcome: | The proposed fingerprint maintains multimodal performance and substantially enhances fingerprint robustness. |
Copied to clipboard
| Challenge: | Recent work shows that Code Large Language Models can address a wide range of code-related tasks. |
| Approach: | They propose a method to generate widespread and versatile instruction data from open source code datasets and use it to train code-related models. |
| Outcome: | The proposed model outperforms open-source models in generalization ability across code-related tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) often perform poorly in generating informative questions, as measured by expected information gain (EIG). |
| Approach: | They propose to use a large language model to enhance the informativeness of LLM-generated questions in 20-question game dialogues by applying a Direct Preference Optimization algorithm to generate low-EIG and high-EI questions. |
| Outcome: | The proposed method produces more effective questions even in domains different from those used to train the DPO model. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable capabilities in machine translation, but most MT-specific LLMs rely heavily on external supervision during training. |
| Approach: | They propose a reinforcement learning framework for machine translation that is reference-free and relies solely on self-judging rewards. |
| Outcome: | The proposed framework outperforms existing LLMs and larger general LLM models on English Chinese translation benchmarks and performs competitively with leading closed-source systems. |
Copied to clipboard
| Challenge: | Existing approaches to augment language models with external knowledge but they are limited by static nature of pre-training data. |
| Approach: | They propose a lightweight approach that compresses retrieved documents into highly dense textual summaries to integrate into in-context RAG. |
| Outcome: | The proposed approach reduces latency and costs while achieving high performance in open-domain questions. |
Copied to clipboard
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios. |
| Approach: | They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios. |
| Outcome: | Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift. |
Copied to clipboard
| Challenge: | Existing open-source vision language models lack high-quality training data for chart reasoning . current models are simplistic and repetitive, while associated QA pairs are prone to hallucinations . |
| Approach: | They propose a framework to synthesize complex charts and reliable reasoning data from scratch. |
| Outcome: | Experimental results show that ChartVerse-8B surpasses existing models in QA and difficulty . lack of high-quality training data hampers development of open-source models . |
Copied to clipboard
| Challenge: | Text-based web agents offer computational efficiency for autonomous web navigation, yet they lack discrimination capabilities to reject plausible but incorrect elements in densely populated pages. |
| Approach: | They propose a model that uses a text-based web agent to learn to discriminate against incorrect elements in densely populated HTML and a training curriculum to synthesize diverse cross-domain tasks with strict verification. |
| Outcome: | Empirical evaluation shows that the model performs better than open-source models with 58.7% step success rate. |
Copied to clipboard
| Challenge: | Existing token-level attacks have shown efficacy on open-source models but suffer from poor cross-model transferability. |
| Approach: | They propose a framework to improve cross-model transferability by modifying model parameters and generating update directions according to differences in output distributions rather than parameter-space distances. |
| Outcome: | The proposed framework improves cross-model transferability and success rates on open-source models. |
Copied to clipboard
| Challenge: | Existing methods to jailbreak Large Language Models (LLMs) exploited internal properties or capabilities of the model, such as optimization-based jailbreak methods and methods that leveraged the model’s context-learning abilities. |
| Approach: | They propose a new method which injects jailbreak information into user prompts and induces the model to generate harmful content. |
| Outcome: | The proposed method achieves near 100% success rates on open-source models while incurring lower time costs compared to previous methods. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. |
| Approach: | They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences. |
| Outcome: | The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences. |
Copied to clipboard
| Challenge: | Large language models (LLMs) struggle to follow complex instructions of IE tasks due to not being aligned with humans. |
| Approach: | They propose an aligned large language moDEL that effectively solves various IE tasks including closed IE, open IE and on-demand IE. |
| Outcome: | The proposed model achieves state-of-the-art (SoTA) performance among open-source models. |
Copied to clipboard
| Challenge: | Existing scientific claim verification benchmarks focus on textual content alone or on verifying claims based on a single table. |
| Approach: | They propose to use SciVer to evaluate the ability of foundation models to verify claims within a multimodal scientific context. |
| Outcome: | The proposed model outperforms 21 state-of-the-art models and human experts on SciVer. |
Copied to clipboard
| Challenge: | Existing methods for benchmarking the uncertainty of large language models face challenges . existing methods require internal model access, additional training, or high computational costs . |
| Approach: | They propose a new benchmark for evaluating the uncertainty of large language models based on confidence intervals . UBench encompasses 11,978 multiple choice questions spanning knowledge, language, understanding, and reasoning capabilities. |
| Outcome: | The proposed method outperforms existing methods for benchmarking the uncertainty of large language models. |
Copied to clipboard
| Challenge: | a capability gap exists between open-source and closed-source large language models (LLMs) . the adoption of closed-sourced LLMs introduces concerns pertaining to openness, privacy, and substantial costs. |
| Approach: | They propose a synthetic data approach that combines strong and weak models for error information . they demonstrate the effectiveness of SENSE, a specialized text-to-SQL model . |
| Outcome: | The proposed method enhances the domain generalization of text-to-SQL models and explores the potential of error data supervision through preference learning. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but the complexity of emerging tasks and higher performance demands highlight the need for continuous improvement. |
| Approach: | They propose a method that refines evaluation results and characterizes model profiles at the knowledge component level. |
| Outcome: | The proposed method improves performance across multiple benchmarks and academic exams. |
Copied to clipboard
| Challenge: | Existing closed-source LLMs have a performance gap in text-to-SQL reasoning tasks. |
| Approach: | They propose a SQL-based approach to synthesize reliable data to enhance text-to-SQL reasoning in LLMs. |
| Outcome: | The proposed model achieves state-of-the-art accuracy on the widely recognized Spider and BIRD benchmarks, significantly narrowing the performance gap with closed-source methods. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable success across diverse tasks such as instruction following, code generation, and medical diagnosis. |
| Approach: | They propose a supervised fine-tuning-based auxiliary loss for Q-value estimations during supervised refinement. |
| Outcome: | The proposed method outperforms beam search on GSM8K, MATH, and GAOKAO on reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to evaluate the capability of large language models to identify lexical semantic equivalence are not currently being used. |
| Approach: | They propose to use the Word-in-Context (WiC) task to determine whether the meanings of a target word remain identical across different contexts to evaluate their capability. |
| Outcome: | The proposed method outperforms other LLMs in the Word-in-Context (WiC) task. |
Copied to clipboard
| Challenge: | Existing models implicitly recover the original text, but it is unclear when they rely on context and when they implicitly do so. |
| Approach: | They propose to use a dictionary to recover adversarial words by using a phonetic, typo, and visual attack to study word recovery performance. |
| Outcome: | The proposed model outperforms open-source models on hateful, offensive, and toxic classification tasks. |
Copied to clipboard
| Challenge: | Recent advances in multi-turn voice interaction models have improved user-model communication, but whether open-source models share this ability remains unexplored. |
| Approach: | They propose to use ContextDialog to evaluate open-source interaction models' ability to recall past utterances to identify key limitations. |
| Outcome: | The proposed model retains and recalls past utterances better than closed-source models, but still struggles with questions about past . findings highlight key limitations in open-source model and suggest ways to improve memory retention and retrieval robustness. |
Copied to clipboard
| Challenge: | relying on proprietary Large Language Models poses privacy and cost implications for models. |
| Approach: | They propose a two-stage fine-tuning approach that breaks down the task into two simpler tasks. |
| Outcome: | The proposed method achieves 60.31% execution accuracy on Bird hold-out test set . it is the highest performance among methods using 7B parameter models . |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for large language models are inflated and inconsistent with actual performance. |
| Approach: | They propose a retrieval-based system to explore potential overlaps between benchmarks and pretraining corpora and a protocol to investigate testset slot guessing. |
| Outcome: | The proposed method exploits overlaps between evaluation benchmarks and pretraining corpora and masks a wrong answer in a multiple choice question and prompts the model to fill in the gap. |
Copied to clipboard
| Challenge: | generative models struggle with logic-intensive instruction following, exposing a persistent reasoning–execution gap. |
| Approach: | They propose a task-agnostic reasoning architecture for general image generation . they propose pixel-level feedback to ground the Thinker's policy in pixel feedback . |
| Outcome: | The proposed system significantly improves image reasoning and generation quality. |
Copied to clipboard
| Challenge: | Recent LLMs have demonstrated promising ability in solving finance related problems, but applying them in real-world finance applications remains challenging due to its high risk and high stakes property. |
| Approach: | They propose a benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. |
| Outcome: | The proposed benchmark outperforms proprietary models in most tasks while open-source models have advantage in specific areas like industry-level fairness. |
Copied to clipboard
| Challenge: | Existing methods for ensembling language models fail to address complex reasoning tasks. |
| Approach: | They propose a framework for process-level ensembling of large language models using Monte Carlo tree search. |
| Outcome: | The proposed framework outperforms both language model decoding and language model ensemble methods on five reasoning benchmarks. |
Copied to clipboard
| Challenge: | Context-DPO is the first alignment method specifically designed to enhance contextfaithfulness for large language models. |
| Approach: | They propose a benchmark that simulates Retrieval-Augmented Generation scenarios with knowledge conflicts to evaluate context-faithfulness. |
| Outcome: | The proposed method improves LLMs' context-faithfulness by 35% to 280% over open-source models. |
Copied to clipboard
| Challenge: | Quantitative reasoning with data is a critical skill to analyze data, yet the assessment of such ability remains limited. |
| Approach: | They propose a quantitative reasoning with data benchmark to evaluate Large Language Models' ability in statistical and causal reasoning with real-world data. |
| Outcome: | The proposed model GPT-4 achieves an accuracy of 58%, while open-source model Deepseek-coder-instruct gets the highest accuracy of 37%. |
Copied to clipboard
| Challenge: | Existing studies focus on detecting machine-generated text in open-source models, but their performance on closed-source large models is limited. |
| Approach: | They propose a method to detect rewritten text from large language models using a BERT encoder and propose to refine it to achieve semantic alignment. |
| Outcome: | The proposed method outperforms baseline methods on three text-generated datasets. |
Copied to clipboard
| Challenge: | Existing issue-resolving frameworks rely on commercial models, leading to high costs and privacy concerns. |
| Approach: | They propose a training approach to enhance issue resolving capability of LLMs by decomposing issue reasolving into subtasks. |
| Outcome: | The proposed approach improves issue-resolving performance and generalizes model . it is cost-effective and provides a cost-efficient alternative to commercial models . |
Copied to clipboard
| Challenge: | Using a dataset of Korean weather queries, we find that automatic speech recognition systems fail on specialized vocabulary. |
| Approach: | They propose an evaluation dataset of Korean weather queries . the dataset was recorded by diverse native speakers following pronunciation guidelines . |
| Outcome: | The proposed model reduces error rates on meteorological terms and improves overall recognition accuracy. |
Copied to clipboard
| Challenge: | Prior studies have shown that large language models can exhibit bias against specific demographic groups and engage in the generation of stereotypical responses. |
| Approach: | They propose a framework to evaluate LLM performance along two axes: safety and utility. |
| Outcome: | The proposed framework evaluates the performance of LLMs along two axes: safety and utility. |
Copied to clipboard
| Challenge: | Solving expert-level multimodal tasks requires strong user query understanding, domain-specific knowledge, and advanced reasoning abilities. |
| Approach: | They propose a benchmark of open-ended user queries encapsulating professional expertise and advanced reasoning. |
| Outcome: | The proposed benchmark is publicly accessible at TBC. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) in Arabic are lacking . despite progress in their development, there is a lack of comprehensive trustworthiness evaluation benchmarks . |
| Approach: | They propose to use Arabic as a language to assess trustworthiness of large language models. |
| Outcome: | The proposed benchmark measures the trustworthiness of large language models in Arabic. |
Copied to clipboard
| Challenge: | figurative language is one of the most challenging aspects of human language for LLMs to comprehend . |
| Approach: | They evaluate LLMs using two multilingual datasets on simile and idiom interpretation and two new evaluation sets for Persian . they find prompt engineering methods are generally effective, but their success varies by figurative type, language, and model. |
| Outcome: | The proposed models perform better in simile and idiom interpretations across languages and figurative types. |
Copied to clipboard
| Challenge: | generative AI has been used to generate fluent and convincing text on social media platforms . a new study examines the generative capabilities of four popular large language models . |
| Approach: | They propose a methodology to examine the generative capabilities of four prominent LLMs on Twitter using a dataset from Llama 3, Mistral, Qwen2 and GPT4o. |
| Outcome: | The proposed method examines the generative capabilities of four prominent LLMs on Twitter. |
Copied to clipboard
| Challenge: | Large language models with instruction-following capabilities are not suitable for long-tail ad hoc extraction use cases for non-expert users. |
| Approach: | They propose a task that follows instructions to extract the desired content from the associated text and present it in a structured tabular format. |
| Outcome: | The proposed paradigm outperforms existing open-source models of similar size in terms of information extraction. |
Copied to clipboard
| Challenge: | Existing methods for large language models (LLMs) use one agent to iterate and execute tools, but they suffer from performance degradation when addressing practical tasks. |
| Approach: | They propose a tool learning framework that coordinates three specialized agents for tool selection, tool execution, and action calibration separately. |
| Outcome: | The proposed framework outperforms baseline models on three datasets with 14% higher success rate. |
Copied to clipboard
| Challenge: | Sycophantic behavior in models can erode user trust by creating a perception of dishonesty or bias. |
| Approach: | They propose to assess the user’s expected answer rather than ignore it and introduce self-augmented preference alignment to reduce sycophancy. |
| Outcome: | The proposed methods significantly reduce sycophancy across tasks and improve models' assessment ability. |
Copied to clipboard
| Challenge: | Recent advances in language models have led to significant improvements in mathematical reasoning across benchmarks. |
| Approach: | They analyze the prevalence of false positives in language models by using heuristic evaluation methods . they find that false positive models produce correct final answers but with flawed deduction paths . |
| Outcome: | The proposed model performance improvements are based on the proposed model and its evaluation metrics. |
Copied to clipboard
| Challenge: | Existing datasets do not cover full range of chart types, such as 3D, volumetric, and gridded charts. |
| Approach: | They propose a hierarchical pipeline and a new dataset for chart generation that leverages the relationships within rich datasets. |
| Outcome: | The proposed method outperforms open-source models and is comparable to state-of-the-art proprietary models in data visualization tasks. |
Copied to clipboard
| Challenge: | Current Vision-Language Models (VLMs) focus on third-person view videos, neglecting the richness of egocentric perceptual experience. |
| Approach: | They propose to use the Egocentric Video Understanding Dataset (EVUD) to train VLMs on video captioning and question answering tasks specific to egocentric videos. |
| Outcome: | The proposed model outperforms open-source models including strong Socratic models using GPT-4 as a planner by 3.6% and outperformed Claude 3 and Gemini Pro Vision 1.0. |
Copied to clipboard
| Challenge: | Using advanced Large Language Models, instructors can improve training of smaller models by analyzing their own model's errors. |
| Approach: | They propose a framework that leverages advanced Large Language Models to enhance training of smaller target models. |
| Outcome: | The proposed framework outperforms ChatGPT on multiple benchmarks and shows that it improves on both in-domain and out-of-domain benchmarks. |
Copied to clipboard
| Challenge: | Currently, large language models are fine-tuned using expensive human-annotated data or GPT-4 generated data. |
| Approach: | They propose to use web-crawled data to train a language model on a smaller set of data . their results show that the model can convert web data with irregular formats into high-quality ones . |
| Outcome: | The proposed model outperforms open-source models larger than 32B and outperformed open-sourced models such as GPT-3.5. |
Copied to clipboard
| Challenge: | Existing models like GPT-3 and Instruct-GPT lack the ability to reformulate unanswerable questions. |
| Approach: | They propose a zero-shot method that combines the strengths of LLMs with a DFS-based algorithm to iteratively explore potential entity combinations and constrain outputs using predefined entities. |
| Outcome: | The proposed method outperforms all baselines, including the GPT-3.5 model, on the unanswerable question reformulation task. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly deployed via third-party system prompts downloaded from public marketplaces. |
| Approach: | They propose a framework that optimizes system prompts to trigger LLMs to output compromised responses only for specific queries. |
| Outcome: | The proposed framework achieves up to 70% F1 reduction on targeted queries with minimal degradation to general capabilities. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are a promising tool for document understanding, but they are not able to handle complex multi-page visual documents. |
| Approach: | They propose a flexible agentic framework for understanding multi-modal, multi-page, and multi-layout documents . SlideAgent employs specialized agents and decomposes reasoning into three specialized levels . |
| Outcome: | a new agentic framework improves accuracy over open-source and proprietary models . it decomposes reasoning into three levels to capture themes and visual cues . the framework is based on a multimodal large language model and a MLLM . |
Copied to clipboard
| Challenge: | Emotion is a central dimension of spoken communication, yet we lack a mechanistic account of how LALMs encode it internally. |
| Approach: | They propose to use emotion-sensitive neurons in large audio-language models to study their interpretations. |
| Outcome: | The proposed models show that they can be used to make decisions on emotion . the results show that the ESNs exhibit non-uniform clustering with partial cross-dataset transfer . |
Copied to clipboard
| Challenge: | Effective interlocutors account for the uncertain goals, beliefs, and emotions of others. |
| Approach: | They propose to calibrate language models to better represent outcome uncertainty . they propose to use two methods to calibrated small open-source models . |
| Outcome: | The proposed fine-tuning strategies can calibrate smaller open-source models to beat pre-trained models 10x their size. |
Copied to clipboard
| Challenge: | Recent advances in generative AI have enabled us to prompt large language models (LLMs) to produce texts which are fluent and grammatical. |
| Approach: | They evaluate model performance by measuring their performance on established benchmarks. |
| Outcome: | The proposed models outperform supervised English GEC models on fluency correction benchmarks and commercial LLMs on edit benchmarks. |
Copied to clipboard
| Challenge: | Image captioning has been a challenge for vision-language researchers for decades . current VLMs focus on tasks like visual question answering (YA) but image captioning is not as advanced as expected. |
| Approach: | They evaluate VLMs' performance on image captioning using human annotations . they find that some metrics show high caption-level agreement with humans . |
| Outcome: | The proposed model outperforms open-source models on image captioning . it achieves 93.4% correlation with human rankings at $4 per test . |
Copied to clipboard
| Challenge: | Multilingual question-answering benchmarks do not factor in regional diversity in the information they capture and tend to be Western-centric. |
| Approach: | They propose to benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics. |
| Outcome: | The proposed model shows greater knowledge of cultural information in English than in the dominant language of the respective culture. |
Copied to clipboard
| Challenge: | OpenCodeInterpreter-33B provides a high level of performance for code generation, executing, and iterative refinement. |
| Approach: | They propose a family of open-source code systems for generating, executing, and iteratively refining code. |
| Outcome: | The OpenCodeInterpreter-33B performs well on humanEval, MBPP, and EvalPlus benchmarks. |
Copied to clipboard
| Challenge: | SWE-Swiss-32B demonstrates strong generalization to other common LLM benchmarks. |
| Approach: | They propose a two-phase training recipe that decomposes issue resolution into three core skills: Localization, Repair, and Unit Test Generation. |
| Outcome: | The proposed model achieves a 60.2% score on the SWE-bench Verified benchmark and is in the top-tier performance bracket of much larger models. |
Copied to clipboard
| Challenge: | Existing Chart2code-related training datasets suffer from limited scale, limited type coverage, and inadequate complexity. |
| Approach: | They propose to synthesize chart2code-related training datasets using web plotting code and chart images to address these challenges. |
| Outcome: | The proposed dataset exhibits the greatest diversity and higher complexity compared to other open-source Chart2code related datasets. |
Copied to clipboard
| Challenge: | Existing models that use Large Language Models (LLMs) show superior performance in various tasks, but lack of controllability leads to unfocused conversations or task failure. |
| Approach: | They propose a standard operating procedure (SOP) framework to regulate dialogue flow by integrating Chain of Thought reasoning and supervised fine-tuning for SOP prediction. |
| Outcome: | The proposed method achieves a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also shows notable gains for open-source models. |
Copied to clipboard
| Challenge: | Existing methods for story evaluation lack reasoning capabilities for open-source models . evolvR framework provides high-fidelity evaluators for story generation tasks . |
| Approach: | They propose a framework that self-synthesizes chain-of-thought data via a multi-persona strategy . they propose evolvR to provide a reward model for story generation . |
| Outcome: | The proposed framework achieves state-of-the-art performance on three evaluation benchmarks . it also enhances the quality of generated stories, validating the superiority of the framework . |
Copied to clipboard
| Challenge: | Current approaches to detect hallucination require many samples from the LLM generator . current methods require multiple samples, which is computationally infeasible . |
| Approach: | They propose a simple baseline for detecting hallucinations in long-form LLM generations . they show that LLM hidden states are highly predictive of factuality in long form natural language generation . |
| Outcome: | The proposed method is comparable to expensive multi-sample approaches while drawing only a single sample from the LLM generator. |
Copied to clipboard
| Challenge: | Prior work on multilingual evaluation has shown that there is a large gap between the performance of Large Language Models on English and other languages. |
| Approach: | They propose to finetune Llama-2 and Mistral models on two datasets to determine their effect on model performance on six downstream tasks covering forty one languages. |
| Outcome: | The proposed model can improve on six multilingual tasks while degrading on high-resource languages. |
Copied to clipboard
| Challenge: | Automated environment configuration is a critical bottleneck in scaling software engineering (SWE) automation. |
| Approach: | They propose a reliable evaluation standard for automated environment configuration for 40 real-world repositories spanning 9 programming languages. |
| Outcome: | The proposed benchmark includes 40 real-world repositories spanning 9 programming languages and measures success in achieving executable states and efficiency under realistic constraints. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks, such as MMLU, C-Eval, and GSM8K, evaluate models by posing a variety of problems, including problems about mathematics, science, law, and general knowledge. |
| Approach: | They propose a benchmark which assesses the model’s lateral thinking within an interactive framework. |
| Outcome: | The proposed evaluation benchmark assesses the model’s lateral thinking within an interactive framework. |
Copied to clipboard
| Challenge: | Existing safety guardrails fail to intercept latent intent, whereas LVLMs can implicitly synthesize holistic malicious semantics from fragmented visual cues. |
| Approach: | They propose an Emoji Chain Hinting Attack (ECHA) framework that decouples sensitive concepts into semantically related emoji chains and structural text masks. |
| Outcome: | The proposed framework outperforms existing baselines and bypasses safety guardrails in over 81% of instances with a single attempt. |
Copied to clipboard
| Challenge: | Existing models exhibit blind tool-use reasoning patterns, which significantly increases inference overhead and degrades model performance. |
| Approach: | They propose an MLLM that performs adaptive tool-use by determining whether a visual problem truly requires tools. |
| Outcome: | The proposed model outperforms existing methods in visual reasoning tasks. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for text-to-audio-video (T2AV) generation are largely designed for human-recorded videos or single-speaker settings. |
| Approach: | They propose a failure-driven diagnostic benchmark for multi-talker dialogue-centric audio-video generation. |
| Outcome: | The benchmark evaluates multi-speaker dialogue generation at four levels: audio-visual signal fidelity, temporal attribute consistency, social interaction, and cinematic expression. |
Copied to clipboard
| Challenge: | Existing video large language models (LMMs) employ an impedance of thousands of frames to understand long videos. |
| Approach: | They propose a plug-and-play module integrated with VideoLLMs to facilitate efficient lengthy video perception. |
| Outcome: | The proposed module boosts the performance of open-source VideoLLMs and proprietary assistants on long-form video benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are popular in the Natural Language Processing community because of their versatility and capability to solve unseen tasks in zero/few-shot settings. |
| Approach: | They investigate the use of large language models in CWI, LCP, and MWE settings by evaluating their use in zero-shot, few-shot and fine-tuning settings. |
| Outcome: | The proposed models struggle in certain conditions or achieve comparable results against existing methods. |
Copied to clipboard
| Challenge: | Recent large language models like ChatGPT and Bard excel in a wide variety of NLP tasks but are not specifically tailored for climate related domain specific information. |
| Approach: | They propose a lightweight Arabic Mini-ClimateGPT that is built on an open-source LLM and specifically fine-tuned on a conversational-style instruction tuning curated Arabic dataset Clima500-Instruct. |
| Outcome: | The proposed model surpasses the baseline LLM in 88.3% of cases during ChatGPT-based evaluation and human expert prefers it over other open-source models. |
Copied to clipboard
| Challenge: | Recent advances in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. |
| Approach: | They propose to redefine the design of vision-language models by identifying key components and creating efficient models with constrained inference costs. |
| Outcome: | The proposed models achieve significant improvements in inference throughput while maintaining high performance. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly ubiquitous, yet their ability to effectively retain and reason about temporal information remains limited. |
| Approach: | They propose six metrics to assess three learning paradigms to enhance temporal knowledge acquisition. |
| Outcome: | The proposed methods improve performance and reduce incorrect outputs. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) are predominantly trained on consistent visual-textual inputs, leaving open the question of whether they can handle semantic mismatches in layout-rich content. |
| Approach: | They propose to use multimodal inconsistency reasoning to assess MLLMs' ability to reason about semantic mismatches in webpages, presentation slides, and posters. |
| Outcome: | The proposed model outperforms open-source models in detecting inconsistencies in webpages, presentation slides, and posters while remaining vulnerable to inconsistent errors. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) have demonstrated impressive performance on existing medical visual question answering benchmarks. |
| Approach: | They evaluate large multimodal models that perform worse than random guessing on medical questions . authors suggest more robust evaluation methods to ensure reliability of LMMs . |
| Outcome: | a new study shows that large multimodal models perform worse than random guessing on medical visual question answering benchmarks. |
Copied to clipboard
| Challenge: | Formalising informal mathematical reasoning into formally verifiable code is a significant challenge for large language models. |
| Approach: | They propose a domain-agnostic human-in-the-loop agentic pipeline to aid autoformalisation in scientific domains. |
| Outcome: | The proposed system produces syntactically correct and semantically aligned proofs for low cost. |
Copied to clipboard
| Challenge: | Existing approaches to creating inclusive vision-language models rely on human annotators, making it labor-intensive and creating cognitive burdens. |
| Approach: | They propose a semi-automated framework for constructing cultural VLM benchmarks . they use an annotated sample of Korean culture to generate questions . |
| Outcome: | The proposed framework is based on a Korean culture dataset and shows that open-source models lag behind proprietary ones in understanding Korean culture. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks focus on single-turn evaluations, overlooking the models’ capabilities in multi-turn interactions. |
| Approach: | They propose a benchmark to evaluate the multi-turn conversational abilities of large language models (LLMs) by analyzing human-LLM conversations and constructing multi-turned queries for each category using GPT-4. |
| Outcome: | The proposed model outperforms open-source models in multi-turn tasks while retaining and recalling historical information. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to remarkable progress across a wide range of natural language processing tasks. |
| Approach: | They propose a training framework that enables fine-tuning LLM agents without human annotation. |
| Outcome: | The proposed framework enables fine-tuning LLM agents without human annotation. |
Copied to clipboard
| Challenge: | Existing benchmarks decompose the end-to-end professional report generation into individual components. |
| Approach: | They propose a benchmarking tool that evaluates 20 real-world professional report generation tasks grounded in multimodal document collections. |
| Outcome: | The proposed model outperforms closed-source models on executive summarization tasks but drops significantly on long-horizon synthesis tasks. |
Copied to clipboard
| Challenge: | We apply definition generators based on open-weights large language models to create explanations of novel senses. |
| Approach: | They apply open-weights large language models to create explanations of novel senses using target word usages as input. |
| Outcome: | The proposed definition generators perform on par with decoder-only models. |
Copied to clipboard
| Challenge: | Existing approaches do not emphasize step-wise problem-solving. |
| Approach: | They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step. |
| Outcome: | The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling. |
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models are limited to multiview diagnostics. |
| Approach: | They propose a benchmark specifically designed for medical multi-image understanding that evaluates MLLMs across four dimensions. |
| Outcome: | The proposed model performs better in multi-image contexts than open-source models . the model perform better when processing increased visual loads than closed-source ones . |
Copied to clipboard
| Challenge: | ComicVQA is a visual reasoning benchmark for comics. |
| Approach: | They propose a comics-based benchmark for evaluating MLLMs on visual reasoning. |
| Outcome: | The proposed model achieves 62.6% accuracy on Missing Panel Prediction and 46.4% on Panel Sorting, compared to open-source models. |
Copied to clipboard
| Challenge: | Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). |
| Approach: | They propose a benchmark for evaluating cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages . they evaluate 12 vision-language models that achieve 70% accuracy when provided with direct context . |
| Outcome: | The proposed benchmark evaluates models with high accuracy over tables and charts extracted from 4,000 Wikipedia pages . proprietary models achieve 70% accuracy when provided with direct context, but open-source models perform worse when retrieval from long documents is required. |
Copied to clipboard
| Challenge: | Existing methods to analyze black-box jailbreaks lack direct optimization signals to refine adversarial prompts. |
| Approach: | They propose a distribution-jailbreak attack method that selects effective jailbreak templates and iteratively optimizes adversarial suffixes by maximizing the KL divergence from the standard refusal distribution. |
| Outcome: | The proposed method achieves state-of-the-art Attack Success Rate (ASR) on all tested open-source models and delivers over 94% ASR on GPT-4.1. |
Copied to clipboard
| Challenge: | AGL1K is the first audio geo-localization benchmark for audio language models (ALMs) it is based on a crowd-sourced platform and is available in 72 countries and territories. |
| Approach: | They propose a benchmark for audio geo-localization that quantifies the informativeness of each recording and a metric that quantizes the information of each audio clip. |
| Outcome: | The proposed benchmarks cover 72 countries and territories and can be used to improve audio geo-localization. |
Copied to clipboard
| Challenge: | Existing video evaluation benchmarks focus on a single language, typically English, and feature videos rooted in Western cultural contexts. |
| Approach: | They propose a video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension. |
| Outcome: | The proposed video evaluation benchmark bridges cultural, linguistic, and domain divides . existing benchmarks only feature videos from YouTube, Shutterstock, or established video datasets based on cultural diversity . |
Copied to clipboard
| Challenge: | Large Language Models struggle with temporal reasoning, which requires processing time-related information such as event sequencing, durations, and inter-temporal relationships. |
| Approach: | They propose a framework that enhances the temporal reasoning abilities of Large Language Models (LLMs) by combining timeline construction with iterative self-reflection. |
| Outcome: | The proposed framework improves the temporal reasoning abilities of large language models and improves traceability of the inference process. |
Copied to clipboard
| Challenge: | Existing prompt-based methods for debiasing are often superficial and lack a thorough understanding of complex bias concepts. |
| Approach: | They analyze a BBQ and stereoSet benchmarks to examine the assumption that large language models understand biases. |
| Outcome: | The proposed model misclassified 90% of unbiased content as biased despite high accuracy on BBQ dataset . the proposed model may have been flawed in previous attempts to debiase . |
Copied to clipboard
| Challenge: | Large language models (LLMs) have remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. |
| Approach: | They propose a framework to train critic models using refinement signals to generate feedback loops where critiques guide the model in refining its responses. |
| Outcome: | The proposed framework outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes. |
Copied to clipboard
| Challenge: | PlotGen-Bench evaluates vision-language models' ability to generate executable visualization code from plots under realistic and complex visualization requirements. |
| Approach: | They propose a benchmark to evaluate plot-to-code generation in vision-language models . they use Matplot, Matplos, Mat3D, Mat4D, and Mat4E to evaluate their performance . |
| Outcome: | The proposed benchmark covers 9 major categories, 30 subcategories, and 3 core tasks . it covers 2D, 3D and animated plots across 5 widely used visualization libraries. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use, but their ability to continuously refine solutions in response to dynamic environmental feedback remains underexplored. |
| Approach: | They propose a benchmark to evaluate self-improvement capabilities in large-scale search spaces by combining 20 machine learning tasks with 10 classic NP-hard problems. |
| Outcome: | The proposed framework emulates human-like cognitive adaptation and operates via a general perception–memory–reasoning loop, iteratively refining solutions based on environmental feedback. |
Copied to clipboard
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
Copied to clipboard
| Challenge: | SafeAgent improves agent safety through fully automated synthetic data generation. |
| Approach: | They propose a framework that improves agent safety through fully automated synthetic data generation. |
| Outcome: | The proposed framework outperforms closed-source models on two safety benchmarks and one real-world task. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names. |
| Approach: | They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings . |
| Outcome: | The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated . |
Copied to clipboard
| Challenge: | Existing approaches to adapt image-focused models for video understanding have not been successful in analyzing long video sequences. |
| Approach: | They propose a video instruction dataset that outperforms existing video instruction data for fine-tuning MLLMs by incrementally increasing input context length. |
| Outcome: | The proposed model outperforms existing models on video benchmarks and outperformed proprietary models on VideoMME even with a compact 7B model. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used as educational tools, yet evaluating their teaching capabilities remains challenging due to the resource-intensive nature of teacher-student interactions. |
| Approach: | They propose a multi-agent dialogue framework that efficiently assesses teaching capabilities through simulated dynamic educational scenarios. |
| Outcome: | The proposed framework outperforms open-source models on 1,498 questions across 13 disciplines and 10 difficulty levels on 1,400 questions. |
Copied to clipboard
| Challenge: | Continual pretraining is an important approach for Large Language Models to improve their performance in target domains, learn new topics and languages, and even boost their general capabilities. |
| Approach: | They propose a training strategy that mitigates instability by increasing the number of epochs, along with two data sampling strategies targeting data domain relevance and corpus distribution. |
| Outcome: | The proposed training strategy improves the average medical task performance of the OpenLlama-3B model from 36.2% to 40.7% using only 40% of the original training budget, while also enhancing general task performance without causing forgetting. |
Copied to clipboard
| Challenge: | VEHME is a vision language model for assessing handwritten math answers . traditional methods of assessing student work are limited by time constraints, class sizes and cognitive load . |
| Approach: | They propose a Vision-Language Model for Evaluating Handwritten Mathematics Expressions to assess handwritten math responses with high accuracy and interpretable reasoning traces. |
| Outcome: | VEHME achieves state-of-the-art performance among open-source models and approaches accuracy of proprietary systems. |
Copied to clipboard
| Challenge: | Existing models that exploit loopholes identify and reason about ambiguity and conflicting goals, presenting a potential safety risk. |
| Approach: | They propose to study the responses of large language models to loopholes by examining ambiguity and pragmatics in LLMs. |
| Outcome: | The proposed models can identify ambiguities and exploit loopholes to satisfy their given goals as opposed to the goals of the user. |
Copied to clipboard
| Challenge: | Clinical Decision Support Systems (CDSSs) provide reasoning and inquiry guidance for physicians, yet they face high maintenance costs and low generalization capability. |
| Approach: | They propose a clinical diagnostic model with clinical reasoning and inquiry skills, the Dr. Assistant, and a pipeline to capture abstract reasoning logic. |
| Outcome: | The proposed model outperforms open-source models and achieves competitive performance to closed-source model. |
Copied to clipboard
| Challenge: | Existing iterative refinement strategies that generate solutions in a single forward pass often hit a performance ceiling on complex algorithmic tasks. |
| Approach: | They propose a reinforcement learning framework that internalizes the structured reasoning trajectory directly into the model’s weights. |
| Outcome: | The proposed framework achieves 94.51% (87.20%) on HumanEval, 81.80% (78.57%) on MBPP, 35.00% on BigCodeBench, 52.21% on LiveCodeBech, and 37.34% on CodeForces in a single-attempt setting. |
Copied to clipboard
| Challenge: | Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation. |
| Approach: | They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. |
| Outcome: | The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models. |
Copied to clipboard
| Challenge: | Podcast script generation is a challenging task for large language models, but evaluation resources are limited. |
| Approach: | They propose a benchmark to evaluate podcast script generation using a multifaceted evaluation framework . PodBench is a prototype that integrates quantitative constraints with LLM-based quality assessment . |
| Outcome: | The proposed framework integrates quantitative constraints with LLM-based quality assessment. |
Copied to clipboard
| Challenge: | Slovak embeddings are core infrastructure for semantic search, retrieval-augmented generation (RAG), clustering, and classification. |
| Approach: | They propose a MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language . they use 31 datasets across 7 task types to evaluate the performance of the models . |
| Outcome: | The proposed model achieves competitive performance with proprietary APIs while remaining locally deployable for RAG . the model is based on 31 datasets across 7 task types and is 4 the depth of existing benchmark for Slovak . |
Copied to clipboard
| Challenge: | Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment. |
| Approach: | They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation. |
| Outcome: | The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are powerful but weak when inputs are perturbed. |
| Approach: | They evaluate LLMs that are more powerful than single LLM in math question answering . they use a unified sampling-and-voting framework to evaluate their models . |
| Outcome: | The proposed models show that collaboration between agents improves accuracy and clean accuracy even with a large number of agents. |